Back

Age and Ageing

Oxford University Press (OUP)

Preprints posted in the last 7 days, ranked by how well they match Age and Ageing's content profile, based on 28 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
People living with multiple long-term conditions have different pathways of unscheduled care in hospital: findings from an analysis of routinely-collected clinical data

Witham, M.; Evison, F.; Bellass, S.; Cooper, R.; Gallier, S.; Pretorius, S.; Sapey, E.; Suklan, J.; Sayer, A. A.

2026-09-01 health informatics 10.64898/2026.08.28.26361696 medRxiv
Top 0.1%
13.4%
Show abstract

Study Objective Little is known about where in hospital care for multiple long-term conditions (MLTC) is delivered. We aimed to describe pathways of care (ward transfers) and outcomes for people admitted to hospital for unscheduled care by MLTC status and other key sociodemographic characteristics. Design and setting Analysis of routinely-collected electronic health records from a large acute UK hospital. Participants Adult unscheduled care admissions from 1st July 2018 to 30th June 2019. The presence of two or more of 59 long-term conditions was ascertained using ICD-10 codes from previous hospital discharges. Main outcome measures Markov state transition probabilities were derived for ward moves and compared for MLTC vs no MLTC, age, sex, ethnicity and neighbourhood deprivation. Outcomes (length of stay, death, readmission, move from definitive ward) and time spent in emergency and assessment departments were compared between subgroups. Results A total of 33,252 adults, mean age 56.0 (SD 21.9) years were analysed; 14,834 (42.4%) had MLTC. People with MLTC were more likely to die in hospital (4.2 vs 1.9%, p<0.001), transfer to internal medicine wards or older peoples medicine wards, were less likely to transfer to surgical wards, had longer median length of stay (1.83 vs 0.69 days, p<0.001), stayed longer in acute medical units (15.5 vs 9.6 hours, p<0.001), and were more likely to move from their definitive ward (18.2 vs 16.4%, p=0.002). Conclusion Unscheduled hospital care pathways are complex and differ for people with MLTC, who have worse outcomes and may be less likely to receive optimal care.

2
The effectiveness of a complex intervention, aimed at reducing hospital occupancy, to improve Emergency Department patient flow: a retrospective controlled interrupted time series

McHenry, R. D.; Caesar, D.; Clarke, B.; Mackay, D.; Pell, J.

2026-09-03 health systems and quality improvement 10.64898/2026.08.31.26361802 medRxiv
Top 0.1%
7.8%
Show abstract

Objectives Emergency department (ED) crowding is recognised as an important public health concern internationally, and is driven principally by exit block, the shortage of inpatient beds for patients requiring admission. This study aimed to evaluate whether a complex intervention targeting hospital occupancy improved ED patient flow, and quantified the change in attendances. Methods A controlled interrupted time series using weekly, publicly reported Public Health Scotland data from 1 January 2022 to 1 February 2026. The multi-component intervention focused on reducing hospital occupancy and included additional adult social care funding; engagement with regional social care providers; accelerated implementation of the Discharge without Delay programme; re-evaluation of whole-hospital escalation thresholds and response; resource and data supporting inpatient department reductions in length of stay; and additional investment in remote clinical assessment. The intervention commenced at a large tertiary ED on 01 February 2025. Primary outcomes were the proportions of attendances spending [&ge;]4, [&ge;]8 and [&ge;]12 hours in the ED. The secondary outcome was attendance volume. Segmented regression was fitted with a contemporaneous control series, seasonal terms and autoregressive moving average errors. Long waits were additionally illustrated as potentially avoided deaths. Results The analysis covered 161 pre-intervention and 52 post-intervention weeks. Relative to pre-intervention levels, the proportion of attendances waiting over 4 hours fell by 10.4% (95% CI 1.6 to 19.2%), by 16.4% (95%CI 1.3 to 31.5%) over 8 hours and by 24.3% (95%CI 2.6 to 46.1%) over 12 hours. Using established associations between long ED waits and excess mortality, by one-year the intervention was potentially associated with 54 fewer excess deaths (95%CI 19 to 93). Attendances rose by 3.8% (95%CI 1.3 to 6.4%) against the counterfactual. Conclusions A complex intervention targeting hospital occupancy was associated with a reduction in long ED waits despite rising attendances. Interventions addressing hospital occupancy can meaningfully improve ED crowding.

3
What Matters Most: A Multi-Stakeholder Study of Outcome Domains in Lower-Limb Prosthesis Use

Ahmed, M. E.; Karlsson-Brown, S.; Koufaki, P.; Ahmadi, M.; Mico-Amigo, E. M.

2026-09-03 rehabilitation medicine and physical therapy 10.64898/2026.08.31.26361544 medRxiv
Top 0.2%
3.4%
Show abstract

Purpose: Lower-limb prosthesis use involves interacting physical, psychosocial, and device-related outcomes that may not be fully captured by conventional clinical assessment. This study aimed to develop and evaluate a stakeholder-informed framework of outcome domains relevant to meaningful everyday prosthesis use. Materials and Methods: A mixed-methods participatory design comprised a structured synthesis of selected clinically relevant content from five established patient-reported outcome measures; semi-structured interviews and importance and actionability ratings with 18 contributors (12 prosthesis users, four clinicians, and two industrial partners); and integration of the synthesis, qualitative, and rating findings. Interview records were analysed using reflexive thematic analysis, and ratings were analysed descriptively. Results: The resulting framework comprised four interrelated domains: Mobility, Physical Function, Psychosocial Wellbeing, and Prosthesis Experience. Mobility showed the clearest convergence across stakeholder perspectives. Prosthesis users showed the largest importance actionability gap for Prosthesis Experience (4.5 vs 3.0), whereas clinicians showed the largest gap for Psychosocial Wellbeing (5.0 vs 3.0). Interviews highlighted day-to-day variability in prosthesis use and the influence of confidence, fatigue, comfort, environmental conditions, social context, and device usability. Conclusions: Meaningful outcome assessment in prosthetic rehabilitation should extend beyond mobility alone to consider physical function, psychosocial wellbeing, and prosthesis experience within everyday contexts. The proposed framework provides a stakeholder-informed foundation for multidimensional outcome assessment in prosthetic rehabilitation.

4
Empowering adults to manage their hearing loss: assessing the benefits of user-controlled, smartphone-connected hearing aids.

Maidment, D. W.; Habib, A.; Gomez, R.; Benton, C.; Ferguson, M. A.

2026-09-03 otolaryngology 10.64898/2026.08.30.26361775 medRxiv
Top 0.2%
1.9%
Show abstract

The availability of hearing aids that can connect wirelessly to smartphone technologies via Bluetooth has grown exponentially in recent years. However, there is limited evidence assessing the benefits of user-adjustability afforded by these devices. This study aimed to assess the benefits of smartphone-connected hearing aids and an accompanying application (or app) in new and existing hearing aid users. In this single-centre, prospective, observational study, 44 adult hearing aid users (14 new and 30 existing) were recruited. Participants were fitted bilaterally with smartphone-connected hearing aids that could be adjusted by the user via an app. Self-reported outcome measures were collected at fitting and after seven-weeks of using the device in everyday life. For both new and existing hearing aid users, significant improvements in social participation, hearing-related fatigue, and hearing aid benefit and satisfaction were found. For existing hearing aid users, all outcomes were significantly better for the smartphone-connected hearing aids plus app in comparison to their existing hearing aids that did not connect to a smartphone, all with moderate-to-large clinical effect sizes (d> .6). User-controllability via the app was identified as the key benefit, and most participants (68%) reported that the app met their needs 'extremely' or 'very well'. These results suggest that, when used in conjunction with an app, smartphone-connected hearing aids can improve hearing outcomes due to greater user-controllability to improve listening. Thus, smartphone-connected hearing aids have the potential to facilitate patient-centred care, empowering the individual to successfully manage their hearing loss.

5
Sleep Intervention for NHS Healthcare Shift Workers: A Pilot Study of Noise-Masking Earbuds

Hickman, R.; Joyce, D. W.; Gray, N.; Shergill, S.; D'Oliveira, T. C.

2026-09-01 psychiatry and clinical psychology 10.64898/2026.08.28.26361673 medRxiv
Top 0.3%
1.2%
Show abstract

Background: Shiftwork disrupts natural sleep-wake cycles, alters light exposure patterns, and contributes to circadian misalignment. Detrimental health consequences associated with shift work include elevated risk for metabolic disorders, cardiovascular disease, cancer and all-cause mortality. Healthcare workers have one of the highest rates of shift work exposure, yet there are relatively few non-pharmacological interventions (with good evidence) developed to improve sleep outcomes in this population. Objective: A pre-post pilot interventional study assessed the acceptability and perceived effectiveness of commercial noise-masking earbuds on improving subjective sleep characteristics among National Health Service (NHS) healthcare staff working fast rotating shifts. Methods: Noise-masking sleep earbuds (Kokoon NightBuds) were worn for a pilot six-week intervention by twenty-seven NHS nurses (aged 26-43 years, 88.9% female) working fast rotating shifts from the EClocker Study. Sensors inside the earbuds were paired with a smartphone app to monitor sleep. An audio library in the smartphone app delivered personalised relaxation exercises and sleep techniques drawn from cognitive behavioural therapy for insomnia (CBT-I). A pre-post two-week monitoring period with daily smartphone-based Experience Sampling Methods (ESM) captured perceived daily sleep patterns. Acceptability and perceived effectiveness of the earbuds in promoting better sleep outcomes was assessed. Results: Use of the noise-masking sleep earbuds over a six-week period was associated with positive sleep improvement trends and elicited promising acceptability. Almost two thirds of NHS fast rotating shift nurses (63%) subjectively reported reductions in general sleep disturbance symptoms (PSQI Global), one in four experienced perceived sleep quality improvements (SQ; 25.9%), one in five reported sleeping longer (TST; 22.2%), and a third perceived falling asleep faster (SOL; 33.3%), had better sleep efficiency (SE; 33.3%) and improved daytime dysfunction (33.3%) (PSQI subcomponent scores). Sleep diaries (CSD) collected daily using smartphone-based ESM also demonstrated small improvements post-sleep earbud use; nurses reported sleeping an average 18 minutes longer (TST) and fell asleep more easily, on average 11 minutes faster (SOL). Sleep earbuds were generally well tolerated; 56% of nurses reported the earbuds as (somewhat to very) helpful, 52% reported (somewhat to strongly) falling asleep more easily (SOL), 44% felt (somewhat to strongly) their sleep quality was improved (SQ) and 30% agreed (somewhat to strongly) they slept longer (TST) and had less disturbed sleep. Conclusions: To our knowledge, this is the first study in Europe to pilot noise-masking earbuds as a potential non-pharmacological aid to improve sleep-wake behaviours or mitigate fatigue for healthcare staff. Preliminary results showed promising acceptability and (small) perceived sleep improvement trends following a targeted six-week earbud intervention in NHS fast rotating shift nurses.

6
Longitudinal Tracking and Construct Validity of a Single-Item Physical Activity Measure in the Womens Healthy Ageing Project

Corcoran, D.; Szoeke, C.; Apostolopoulos, V.; Feehan, J.

2026-08-31 public and global health 10.64898/2026.08.27.26361567 medRxiv
Top 0.3%
1.2%
Show abstract

This study aimed to quantify the longitudinal tracking and cross-sectional construct validity of a single-item questionnaire measuring recreational physical activity frequency (RPAF) in the Womens Healthy Ageing Project. At baseline, 474 participants aged 45-55 reported RPAF from 1993 to 2014. Longitudinal tracking of the RPAF item was assessed as a consecutive-wave and baseline-referenced measure using linear weighted kappa (LWK), Spearman correlations, exact agreement and within-one-category agreement. Construct validity in the form of convergent and known-group validity was assessed using the International Physical Activity Questionnaire (IPAQ) leisure activity domains, Short Form 36 physical function (SF-36-PF) subscale, Timed Up and Go (TUG), hand grip strength (HGS) and waist-to-height ratio (WHtR). 474 participants provided baseline RPAF data. Pairwise longitudinal samples ranged from 176 to 459 across the study. Consecutive-wave LWK ranged from 0.38 to 0.49, and Spearman correlations ranged from 0.44 to 0.57. Exact and within-category agreement ranged from 41.4%-50.8% and 72.0%-79.0%. Baseline-referenced LWK ranged from 0.22 to 0.47, with Spearman correlations of 0.29 to 0.56. RPAF correlated with total IPAQ leisure score (rs = 0.60), IPAQ walking score (rs = 0.58), SF-36-PF (rs = 0.33) and TUG score (rs = -0.25). No significant correlation was identified between RPAF, HGS or WhTR. RPAF discriminated known groups for WHO guideline-sufficient activity, SF-36-PF, and TUG fall risk. The RPAF item demonstrated fair-to-moderate agreement in consecutive waves, with weaker baseline-referenced tracking. Cross-sectional validity was highest with total IPAQ leisure activity. The item may provide a pragmatic measure for RPAF in womens cohort studies.

7
Temporal Dynamics of Daily Sleep, Mood, and Cognition in NHS Shift Workers: A Digital Experience Sampling (ESM) Study

Hickman, R.; Joyce, D. W.; Gray, N.; Hampshire, A.; Hellyer, P. J.; Cai, Z.; Shergill, S.; D'Oliveira, T. C.

2026-08-31 psychiatry and clinical psychology 10.64898/2026.08.27.26361516 medRxiv
Top 0.3%
1.1%
Show abstract

Background Sleep, mood, and affective states are mutually connected. There is a paucity of studies, however, that have considered bidirectional relationships between daily sleep-affective dyads in naturalistic settings, particularly for shift workers. Objective To evaluate the dynamic and temporal interplay of daily smartphone-based self-reported sleep measurements, dimensions of affective experience and cognitive processing in UK shift working nurses. Methods The EClocker Study prospectively monitored 102 National Health Service (NHS) nurses (aged 25-61 years, 83.3% female) working standard (day shift) and non-standard (fast rotating shifts) schedules over a two-week period. Smartphone-based Experience Sampling Methodology (ESM) recorded daily sleep, mood, momentary affect and cognitive attentional functioning. Self-reported burnout, emotional dysregulation, emotion reactivity and affective dimensions (positive and negative) were also collected. Findings Overall, NHS nurses reported a high prevalence of depressive symptoms, stress, burnout and sleep-circadian rhythm disturbances. Generalised Additive Modelling (GAMs) revealed that NHS nurses higher perceived sleep quality predicted better next-day mood state, while better daytime mood was associated with reduced sleep onset latency, such that participants reported falling asleep faster. In contrast, daytime mood or affect (positive and negative) had no substantial, direct impact on nurses subjective sleep parameters (sleep quality, sleep duration, sleep efficiency). Exposure to fast rotating night shifts across the two-week study was associated with more frequent response errors on a Choice Reaction Time (CRT) cognitive task, while daytime somnolence did not adversely influence nurses momentary reaction time speeds or attentional function. Conclusions Clinically relevant sleep impairments, insomnia-related symptoms, elevated stress, and poor mood were pervasive in a sample of UK NHS nurses, regardless of shift type. Sleep quality impacted next-day mood and daytime mood impacted sleep latency, while rotating shifts led to an increase in cognitive errors. Recognising the impact of shiftwork and designing interventions to promote better sleep quality offer potential to enhance mood and performance in healthcare professionals. Clinical implications We need to implement and evaluate interventions that regularise sleep patterns and promote sleep quality to alleviate mood symptoms among frontline NHS shift workers.

8
Evaluating Clinical Foundation Models for Early Alzheimer's Disease and Related Dementia Prediction from Longitudinal EHRs

Farzana, S.; Arian, A.; Rundek, T.; Desvarieux, M.; Ahsan, H.

2026-09-03 health informatics 10.64898/2026.09.01.26361933 medRxiv
Top 0.4%
1.1%
Show abstract

Early identification of Alzheimer's disease and related dementias (ADRD) remains challenging despite its importance for timely intervention, management of modifiable risk factors, and care planning. We developed and evaluated ADRD onset prediction models using longitudinal electronic health records (EHRs) from the All of Us Research Program at clinically meaningful lead times of 6, 12, 24, and 36 months before diagnosis, benchmarking interpretable count-based representations against four publicly available pretrained clinical foundation models (CLMBR-T, GPT-style, LLaMA-style, and Mamba) across multiple ADRD phenotype definitions. Count-based models consistently achieved the highest discrimination and calibration across all cohorts and prediction horizons. Predictive performance declined with increasing lead time for all approaches; however, the performance gap between count-based and pretrained representations progressively narrowed, with foundation models achieving comparable AUROC of 0.719 (compared to the AUROC of 0.738 of count-based model) at the 36-month horizon while providing higher sensitivity and F1 scores under a fixed operating threshold. External validation with zero-shot evaluation on UChicago EHRs exhibited limited generalizability for count-based and pretrained clinical foundation model based representations. These findings demonstrate that transparent count-based EHR representations remain the strongest overall approach for ADRD onset prediction, while pretrained clinical foundation models provide complementary advantages for long-term risk identification and establish a benchmark for evaluating transferable clinical representations in temporal ADRD risk prediction.

9
Artificial Scientific Intelligence for Measurement-burden-aware Modelling and Interpretation of Multi-site Bone Mineral Density

Xiang, S.; He, H.; Xie, Z.; Cheng, C.-Y.; Li, H.; Liu, D.

2026-09-01 health informatics 10.64898/2026.08.30.26361665 medRxiv
Top 0.4%
1.0%
Show abstract

Agentic workflows can coordinate modelling, but balancing predictive performance, measurement burden and reproducibility is unclear. We developed DXA Agent, an agentic workflow for dual-energy X-ray absorptiometry (DXA) outcomes integrating planning, feature-model refinement, tools, provenance and hypothesis-generating interpretation. Models were independently developed and tested in UK Biobank (5,318 participants) and the National Health and Nutrition Examination Survey (NHANES; 3,777 participants), using cost-efficient and no-limit strategies. Across 20 UK Biobank and three NHANES bone mineral density sites, cost-efficient models achieved lower RMSE and higher R2 than the best conventional comparator, with median relative RMSE reductions of 10.9% and 9.9%, respectively. Classification was task dependent: UK Biobank osteoporosis averaged AUROC 0.839 and PR-AUC 0.182, whereas NHANES performance was comparable with conventional models. Higher-burden features did not consistently improve prediction. These retrospective, cohort-internal findings position DXA Agent as an inspectable, measurement-burden-aware research workflow requiring independent prospective validation.

10
External Validation of a Mathematical Model of Brain Health

Sadia, H.; Doyon, N.; Duchesne, S.

2026-09-03 neurology 10.64898/2026.09.01.26361929 medRxiv
Top 0.4%
0.9%
Show abstract

Background Understanding the mechanisms underlying brain aging and age-related pathological changes is essential for advancing brain health research. Our group previously developed a mechanistic mathematical model of healthy brain, Chamberland et al. (2024) that integrates key biological processes involved in normal aging, from which Alzheimer's disease (AD) related changes may emerge naturally. Objectives To characterize and validate this brain model by evaluating its sensitivity, calibrating its parameters, and assessing generalizability in independent populations. Methods The model represents the evolution of key biological processes associated with brain aging, including amyloid beta (A{beta}), tau pathologies, neuroinflammation, and neuronal death. After identifying the 30 most influential parameters, we calibrated the model using cognitively normal (CN) participants from the AD Neuroimaging Initiative (ADNI) database (n = 211) by minimizing a loss function composed of three outcomes (AB) plaques, tau tangles, and neuronal density). The calibrated model was then applied to the UK Biobank cohort (n = 35,899) of normal controls (aged 44-82 years). The effects of sex and APOE were evaluated using stratified simulations. Results Parameter calibration significantly reduced the prediction errors for A{beta} and tau. Neuronal density predictions showed strong agreement in the UK Biobank cohort. The variance decomposition identified APOE status as a major contributor to variability in A{beta}. Conclusion Our validated brain health model links mechanistic pathways with population data and reproduces neuronal density patterns in an independent cohort. These findings support its use as a framework for studying brain aging and investigating how Alzheimer's disease related pathological changes may emerge with aging.

11
Default-filled outcome labels in a deployed cognitive-screening programme: an operator-level audit and the construction of twenty-four language-model arms

Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.

2026-09-02 health informatics 10.64898/2026.08.28.26361585 medRxiv
Top 0.5%
0.8%
Show abstract

Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.

12
Long-term outcomes of cruciate ligament injury: evidence from New Zealand linked register data

Pryymachenko, Y.; Wilson, R.; Abbott, J. H.

2026-09-01 epidemiology 10.64898/2026.08.27.26361565 medRxiv
Top 0.6%
0.5%
Show abstract

Objectives To analyse the long-term effects of a cruciate ligament (CL) injury on health and socioeconomic outcomes. Methods We used a comprehensive national injury insurance database to identify CL injuries occurring in New Zealand between 2009 and 2022, and employed a doubly robust staggered difference-in-differences research design to identify the effects of these injuries on outcomes up to 10 years after injury. The outcomes of interest were healthcare use (hospitalisations, emergency department visits, medications, knee replacement surgery for osteoarthritis), associated healthcare costs, and labour market outcomes (employment rates, income, and government benefit payments). Results We identified 61 344 CL injuries for inclusion in the analysis. Over 10-year follow-up, a CL injury resulted in increased healthcare use (0.6 more hospitalizations [95%CI 0.4 to 0.7], 1.7 more days spent in hospital [95%CI 1.3 to 2.1], 0.4 more emergency department visits [95%CI 0.3 to 0.6], 2.5 more outpatient visits [95%CI 1.8 to 3.2], and 4.7 more medications dispensed [95%CI -1.8 to 11.2]) and public healthcare costs ($7 537; 95%CI 5 888 to 9 186), reduced income (-$6 060; 95%CI -11 644 to -475), and increased benefit payments ($1 152; 95%CI 542 to 1 761). Conclusion CL injuries have long-term impacts on healthcare use and socioeconomic outcomes. Strategies to reduce the incidence of CL injuries have the potential to realise large health and economic benefits.

13
Returning APOE and pTau-217 Results: the eSMARTER Randomized Noninferiority Clinical Trial

Langbaum, J. B.; Erickson, C. M.; Langlois, C.; Wood, E. M.; Egleston, B. L.; Harkins, K.; Mim, R.; John, S.; Brown, C.; Brown, S.; Howe, S.; Cacioppo, C.; Eppelmann, L.; Enos, J.; Salata, H.; DeSantiago, D.; Largent, E. A.; Reiman, E. M.; Denkinger, M. N.; Ashton, N. J.; Roberts, J. S.; Karlawish, J.; Bradbury, A. R.

2026-09-01 neurology 10.64898/2026.08.27.26361535 medRxiv
Top 0.6%
0.5%
Show abstract

Importance: Patients are increasingly learning Alzheimers disease (AD) genetic and biomarker results through electronic health portals. Evaluation of alternative scalable delivery models for return of AD risk information is needed to best support patient understanding and psychological well-being. Objective: To determine whether a patient-centered digital platform is comparable to clinician-mediated telehealth sessions for returning APOE and plasma pTau-217 results on outcomes of knowledge and psychological well-being. Design: The Evaluation of Self-Mediated Alternatives for Risk Testing Education and Return of Results (eSMARTER) study was a noninferiority trial of a patient-centered digital platform compared to clinician-mediated disclosure of APOE genotype and optional pTau-217 disclosure. Setting: Decentralized, fully remote trial enrolled participants in the contiguous United States (U.S.) between October 2024 and February 2025, with follow-up completed in November 2025. Participants: Eligible participants were aged 60-80 and had previously undergone APOE genotyping (without disclosure) via the GeneMatch program, passed psychological screening, had internet access, and were English-speaking. Interventions: Participants were randomized, 2:1, to the eSMARTER digital platform or clinician-mediated disclosure of APOE genotype. Following the 6-month post-APOE assessment, participants were offered optional pTau-217 disclosure via the same randomized modality. Main Outcomes and Measures: Primary outcomes at 1-7 days following APOE disclosure included changes in anxiety, disease-specific distress, and AD-related knowledge within a priori non-inferiority margins. Results: 674 persons (mean [SD] age 68 [4.7] years; 451 [67%] female; mean [SD] telephone MoCA=19 [2]) were eligible and provided demographic information. 651 participants were randomized to clinician-mediated (n=216) or digital disclosure (n=435) and completed APOE disclosure (66 [10%] APOE4 homozygotes, 377 [58%] heterozygotes, 208 [32%] non-carriers). 604 participants completed the study; 500 completed optional pTau-217 disclosure. Baseline characteristics were balanced across groups. At 1-7 days following APOE disclosure, scores on AD-related knowledge, PROMIS Anxiety, and disease-specific distress measures met non-inferiority. Conclusions and Relevance: Disclosure of APOE genotype by the eSMARTER digital platform is non-inferior to clinician-mediated telehealth disclosure. No significant between group differences were found following disclosure of pTau-217 results. Together, these results suggest that this digital platform may provide an evidence-based scalable approach for returning AD genetic and biomarker results.

14
The effect of high-dose glucocorticoids on opioid consumption in the first 24 hours after elective hip and knee arthroplasty: A natural experiment study of 47,317 surgeries in Eastern Denmark

Laigaard, J.; Moeller, M. O.; Olsen, M. H.; Overgaard, S.; Mathiesen, O.; Karlsen, A. P. H.

2026-09-02 pain medicine 10.64898/2026.08.31.26361793 medRxiv
Top 0.7%
0.3%
Show abstract

Background: In Denmark, perioperative high-dose glucocorticoid treatment were step-wisely implemented for total hip arthroplasty (THA), total knee arthroplasty (TKA), and unicompartmental knee arthroplasty (UKA). We aimed to estimate the effect of a single high dose of glucocorticoids on opioid consumption following primary THA, TKA, and UKA. Methods: This was a prespecified analysis of a multicenter natural experiment using electronic health record data. We included elective THA, TKA, or UKA surgeries performed in Eastern Denmark from 2017-2025. At each center, surgeries before implementation of high-dose glucocorticoids served as controls, whereas surgeries after implementation comprised the intervention group. The primary outcome was the between-group difference in cumulative 0-24h opioid consumption, which included preemptive end-of-surgery doses. The predefined minimal important difference was set at 5 mg IV morphine equivalents. Secondary outcomes were maximum 0-10 numerical rating scale (NRS) pain score and incidence of opioid-related adverse events within 24 hours, hospital length of stay, and days alive and out of hospital at 30 days. Results: A total of 47,317 surgeries performed at nine centers were analyzed: 13,010 controls and 34,307 in the intervention group. During the study period, five centers implemented high-dose glucocorticoids for THA patients, two for TKA/UKA patients. High-dose glucocorticoids were administered to 6% of patients before implementation versus 92% after. High-dose glucocorticoids resulted in a mean reduction of 3.8 mg intravenous (IV) morphine equivalents (95% CI 3.3;4.3). The intervention also reduced the maximum 0-24h NRS pain score by 0.8 points (99% CI 0.7;0.9), but there was no difference in adverse events, length of stay, or days alive and out of hospital. Conclusions: Implementation of high-dose glucocorticoids reduced 0-24-hour opioid consumption by 3.8 mg IV morphine equivalents after elective hip and knee arthroplasty. This difference was below the prespecified minimal important difference threshold. Online registration: https://doi.org/10.1101/2025.11.11.25339982

15
The implementation of an unscheduled care co-ordination hub (Flow Navigation Centre Plus), and emergency department attendances and delays: a controlled interrupted time series.

McHenry, R. D.; Moultrie, C. E.

2026-08-31 emergency medicine 10.64898/2026.08.28.26361651 medRxiv
Top 0.7%
0.3%
Show abstract

Objectives Emergency Department (ED) crowding is an international concern, predominantly caused by 'exit block', the lack of availability of inpatient beds for those requiring admission. The implementation of Flow Navigation Centre Plus (FNC+) services in Scotland aimed to reduce self-presentation to EDs and reduce crowding by providing remote clinical assessment for patients contacting urgent care by telephone and professional-to-professional advice on patient pathways, but their effectiveness is unknown. This study aimed to estimate the effect of board-wide implementation of FNC+ on ED attendances and long waits during the first year of FNC+ operation. Methods Controlled interrupted time series using weekly, publicly reported Public Health Scotland data. The intervention was implementation of the FNC+ in NHS Lanarkshire on 1 April 2024. Counts were summed across constituent sites and percentages derived from board totals. Co-primary outcomes were ED attendance volume and the proportions of attendances spending more than 4, 8 and 12 hours in the department. Segmented regression was fitted with contemporaneous control boards, seasonal terms, and accounted for autoregression. Results 118 pre-intervention and 52 post-intervention weeks were analysed across all 3 EDs in the implementing board. Attendances showed no detectable step change (+1.20%; 95%CIs -0.66 to +3.10) relative to the counterfactual. The estimated effect increased across follow-up, however, changing by +3.95% over 52 weeks (95% CI +0.36 to +7.67%). There was no significant step change in the proportion of attendances waiting more than 4 hours following the intervention (+1.74%; 95%CIs -0.71 to 4.20%). Some transition and structural sensitivity analyses demonstrated significant deteriorations in ED performance, and increased attendances, in the year following implementation, and none demonstrated improvements. Conclusions Board-wide implementation of a Flow Navigation Centre Plus was not associated with a step change in ED attendances or in long waits, but there is some evidence that attendances increased and long waits increased in the year following implementation. Their provision of supply-sensitive care is a possible mechanism. Additionally, given their action at the point of input, aiming to divert patients from ED attendance, it is unlikely that such services could relieve a constraint due to exit block, the availability of inpatient care for those requiring admission.

16
Causal roles of phenotypic age and metabolic health on dementia: a Mendelian randomisation and structure learning study

Baousi, A.; Dobinda, K.; Zhu, J.; Yu, X.; Muir, K.; Lophatananon, A.; McMillan, B.; Clarkson, P.; Tang, E. Y. H.; Guo, H.

2026-09-03 genetic and genomic medicine 10.64898/2026.09.01.26360731 medRxiv
Top 0.7%
0.3%
Show abstract

Background Phenotypic age acceleration (PhenoAgeAccel), derived from PhenoAge, and MetaboHealth are composite exposures of biological ageing and metabolic health associated with dementia-related outcomes. Whether these associations are causal and reflect the exposures, constituent biomarkers, or both remains unclear. Methods This study included UK Biobank participants of White British genetic ancestry. MetaboHealth was derived from nuclear magnetic resonance (NMR) metabolomics and PhenoAgeAccel from clinical biomarkers and chronological age. Genome-wide association studies (GWAS) were conducted for MetaboHealth (n=272,568) and PhenoAgeAccel (n=274,077). Independent genome-wide significant variants were used as genetic instruments in two-sample Mendelian randomisation (MR) with FinnGen all-cause dementia summary statistics. Inverse-variance weighting was the primary MR method. Causal network analysis estimated relationships among constituent biomarkers and dementia. Findings GWAS identified 126 and 141 independent genome-wide significant variants for MetaboHealth and PhenoAgeAccel, of which 109 and 141 were retained as genetic instruments. MR found no evidence of a causal effect of genetically predicted MetaboHealth (per unit: OR 0.83, 95% CI 0.49-1.42; p=0.51) or PhenoAgeAccel (per year: OR 0.99, 95% CI 0.95-1.02; p=0.44) on all-cause dementia, with consistent findings across sensitivity analyses and robust MR methods. Lower lymphocyte percentage and higher NMR-derived glucose had direct relationships with dementia in the joint constituent-biomarker network. Interpretation MR provided no evidence that either composite exposure causally influenced dementia. The network prioritised lymphocyte percentage and NMR-derived glucose, supporting examination of composite exposures alongside their constituent biomarkers. Funding NIHR, UKRI, MRC, UK Dementia Research Institute, Innovate UK, and European Union. Full funding details are provided in the acknowledgements.

17
Are Frontier Large Language Models Safer Than Government-Backed Symptom Checkers for Clinical Self-Triage? A Standardised Vignette Evaluation

Chowdhury, A. R.; Chowdhury, B.

2026-09-02 health informatics 10.64898/2026.09.01.26361908 medRxiv
Top 0.8%
0.3%
Show abstract

Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.

18
Acute Renal, Hepatic, Thromboembolic and Functional Complications after Community-Acquired Acute Lower Respiratory Tract Infection: A Prospective Cohort Study in Bristol, UK, 2022-2024

Chatzilena, A.; Hyams, C.; Challen, R.; Lahuerta, M.; McGuinness, S.; Clout, M.; Begier, E.; King, J.; Morales-Aza, B.; Duale, K.; Rodriguez Pereira, A.; Healy, W.; Southern, J.; Wells, P.; Lihou, K.; Grimes, C.; Campling, J. A.; Maskell, N.; Oliver, J.; Vyse, A.; Gessner, B.; Finn, A.; Danon, L.; The AvonCAP Research Group,

2026-09-02 respiratory medicine 10.64898/2026.08.28.26361617 medRxiv
Top 0.8%
0.3%
Show abstract

Introduction Acute lower respiratory tract disease (aLRTD) is a leading cause of hospitalisation and death, particularly in older adults and adults with comorbidities, with acute lower respiratory tract infection (aLRTI; pneumonia and non-pneumonic LRTI) being a major component. Non-pulmonary complications and functional decline after aLRTI are recognised, but their pathogen-specific burden is poorly described. We aimed to quantify renal, hepatic, thromboembolic and functional complications, and mortality, after aLRTI hospitalisation, by clinical phenotype and pathogen. Methods We conducted a cohort study of adults (>18 years) admitted with aLRTD to two hospitals in Bristol, UK (01 August 2022-31 July 2024). aLRTD was classified as pneumonia, non-pneumonic LRTI (NP-LRTI) or no diagnosis of aLRTI. Pathogens were identified from standard-of-care and research microbiology. Outcomes were acute kidney injury (AKI), acute liver dysfunction, venous thromboembolism (VTE), in-hospital falls, reduced mobility at discharge, increased care requirements, and 30-day and 1-year mortality. Analyses were descriptive. Results Among 246,797 adult admissions, 21,456 aLRTD hospitalisations were included: 10,239 (47.7%) pneumonia, 7,742 (36.1%) NP-LRTI and 3,475 (16.2%) with no evidence of aLRTI. Of 19,152 tested aLRTD admissions, 8,503 (44.4%) had a positive microbiological/virological test, yielding 9,204 pathogen detections; 1,194 (6.2%) had co-infections, and SARS-CoV-2 was most frequent, with influenza the second most common in pneumonia and NP-LRTI. Pneumonia had greater severity than NP-LRTI and no diagnosis of aLRTI (median length of stay 6 vs 4 vs 4 days; ICU admission 3.4% vs 0.7% vs 0.5%, respectively). Overall, 22.2% developed AKI, 6.1% acute liver dysfunction, 0.6% DVT and 2.4% PE; 1.8% had a fall, 11.5% reduced mobility, and 16.6% required increased care at discharge. 30-day and 1-year mortality were highest for pneumonia (14.0% and 32.0%, respectively). Pathogen-specific analyses showed longer stays and higher complications and mortality rates for SARS-CoV-2 and Streptococcus pneumoniae, and shorter stays with lower complication and mortality rates for influenza and Haemophilus influenzae. Conclusions Non-cardiovascular complications and functional decline after aLRTI were common, particularly in pneumonic and SARS-CoV-2 or pneumococcal disease. These findings support routine surveillance for renal, hepatic, thromboembolic events, early mobilisation and rehabilitation, and consideration of multi-system outcomes when evaluating public health and economic value of vaccines and therapies.

19
Half of alcohol, drug, and self-harm presentations cannot be identified in coded emergency department data: a diagnostic accuracy study of a large language model

Humphries, C.; Brett, J.; Gruber, F.; James, E.; McKendrick, T. I.; McNairn, K. C.; Miell, A.; O'Brien, R.; Rahman, F.; Schölin, L.; Stewart, M.; Casey, A.

2026-08-31 health informatics 10.64898/2026.08.26.26361443 medRxiv
Top 0.9%
0.2%
Show abstract

Objective To measure the accuracy of clinical coding, clinician review, and a locally deployed large language model (LLM) in identifying alcohol, drug, and self-harm involvement in emergency department (ED) attendances, and quantify prevalence. Design Two-phase diagnostic accuracy study. In a validation week, the identification strategies were assessed against a conflict-adjudicated reference standard (n=2,256); the LLM was then applied to n=105,096 annual attendances at the same site. Setting UK Type 1 Emergency Department treating patients [&ge;]16yrs. Main outcome measures Prevalence quantification compared with the reference standard; sensitivity, specificity, and balanced accuracy of each strategy; monthly identification rates and adjusted annual prevalence. Results The reference standard identified 12.1% of attendances as involving alcohol, drugs, or self-harm (coding 6.0%; clinician 10.0%, LLM 15.6%). LLM balanced accuracy matched or outperformed clinician review in all three domains (alcohol 0.942 v 0.930, p=0.635; drug 0.959 v 0.791, p<0.001; self-harm 0.982 v 0.908, p=0.004). Coding recorded 1.07 domains per identified patient against 1.32 in the reference standard. Adjusted annual prevalence corresponded to 12,890 domain involvements per year not identifiable in coded data. Subdomain classification found at least 81.6% of self-harm attendances required medical assessment for injury or overdose before psychiatric review. Conclusions Clinical coding identified fewer than half of presentations involving alcohol, drugs, and self-harm and rarely captured co-occurring domains; under-recording was present across a full year. A locally deployed LLM generated more complete structured data from existing clinical text within NHS infrastructure, at a scale which is not feasible for manual review.

20
From Bone-centric to Kidney-centric: Environment-Dependent Shift of Spaceflight Renal Stone Pathways

Shi, J.; Gu, Q.; Pan, J.; Yang, A.; Fan, M.

2026-08-31 urology 10.64898/2026.08.27.26360881 medRxiv
Top 1%
0.1%
Show abstract

Human deep-space missions face bone-kidney risks that cannot be extrapolated from six-month ISS data. We built a 12-state Ca-bone-urine-stone mechanistic ODE model and jointly calibrated its 11 physiological parameters on eight ISS targets by Bayesian identification (M0 base = 19-D; M1 extension adds a GCR-bone coupling term for parsimony testing only), then propagated the M0 posterior to four environments (ISS, Lunar subsurface, Lunar surface, Mars). Lumbar-lower BMD loss increases with mission duration and partial-gravity unloading (ISS 180 d -4.83% -> Mars 730 d -12.15%; 2^3 factorial: duration 82.9%, gravity 12.5%, GCR main effect ~ 0), whereas stone rate follows the opposite gradient (ISS 16.1 vs Mars 13.1 per 1000 person-years), reflecting weakened partial-gravity bone resorption alongside residual urinary chemistry changes. The dominant pathway thus shifts from bone-centric on the ISS to kidney-centric on Mars, where residual urinary-chemistry changes-not bone resorption-drive stone risk. The direct GCR-bone coupling term is unidentifiable at current ISS doses (DeltaWAIC = +0.0076 +/- 0.126 SE), so M0 is retained as the main inference model. Bisphosphonates provide >=84% BMD protection but leave a urinary-chemistry residual, so bisphosphonate monotherapy would underestimate Mars stone risk; potassium-magnesium-citrate combinations (RRR_RSS 51%) should therefore be added to deep-space countermeasures. A Lunar-surface 365-day mission is the earliest environment on the NASA roadmap to cross a composite RED threshold. That profile differs from the regolith-shielded 180-day case in both cumulative GCR (~69x) and duration (2x), so a shielding-specific effect cannot be isolated here; forcing the GCR coupling terms to zero leaves all four composite tiers unchanged (0/4, Supp S24), and the shielded 180-day profile is YELLOW rather than GREEN. Independent hold-out validation (Culliton 2025 60-day HDT-bedrest RCT, n=8 control arm of n=24 total) supports the M0 posterior predictive distribution on the lumbar-BMD sub-scope.